Papers with hypothesis generation

16 papers
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization (2026.acl-srw)

Copied to clipboard

Challenge: Automatic prompt optimization (APO) is a powerful paradigm for improving LLM performance without manual prompt engineering.
Approach: They propose a framework that decouples hypothesis generation from prompt rewriting . they propose VISTA framework that recovers accuracy to 87.57% on same defective seed .
Outcome: The proposed framework outperforms baselines on GSM8K and AIME2025 on a defective seed.
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)

Copied to clipboard

Challenge: EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences.
Approach: They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Outcome: EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences.
Literature Meets Data: A Synergistic Approach to Hypothesis Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for hypothesis generation are theory-driven and data-driven, but they lack the computational power to complement each other.
Approach: They develop a method that combines literature-based insights with data to perform LLM-powered hypothesis generation.
Outcome: The proposed method outperforms baseline methods on five datasets and shows human accuracy improves on deception detection and AI generated content detection tasks.
LMdiff: A Visual Diff Tool to Compare Language Models (2021.emnlp-demo)

Copied to clipboard

Challenge: LMdiff visually compares probability distributions of two different language models . notably absent from the range of available tools are those that aim to compare distributions produced by different models.
Approach: They propose a tool that visually compares probability distributions of two different language models that differ through finetuning, distillation, or simply training with different parameter sizes.
Outcome: The proposed tool allows the generation of hypotheses about model behavior by investigating text instances token by token and further assists in choosing interesting text instances from large corpora.
DIXITWORLD: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay (2026.acl-short)

Copied to clipboard

Challenge: Existing evaluations of multimodal abductive reasoning are limited to static, single-agent tasks.
Approach: They propose a multiagent evaluation suite that deconstructs the current evaluations of multimodal abductive reasoning in vision–language models.
Outcome: The evaluation suite is based on two core components: DixitArena and DixitsBench.
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation (2024.acl-long)

Copied to clipboard

Challenge: Abductive reasoning is the process of making educated guesses to provide explanations for observations.
Approach: They propose a task of complex logical hypothesis generation to generate a complex logique hypothesis that can explain a set of observations.
Outcome: The proposed model generates logical hypotheses closer to the reference hypothesis, but not better on unseen observations.
EA2E: Improving Consistency with Event Awareness for Document-Level Argument Extraction (2022.findings-naacl)

Copied to clipboard

Challenge: Recent work on document-level event argument extraction models each individual event in isolation and therefore causes inconsistency among extracted arguments across events.
Approach: They propose an event-aware argument extraction model with augmented context to improve consistency . they hypothesize that participants tend to play consistent roles across multiple events in a document .
Outcome: The proposed model improves consistency and accuracy of arguments extracted from documents.
On the Role of Model Prior in Real-World Inductive Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have evaluated the inductive reasoning capabilities of Large Language Models (LLMs) by evaluating their ability to generate textual hypotheses based on in-context input-output pairs and test these hypothese based upon unseen examples.
Approach: They evaluated three inductive reasoning strategies across five real-world tasks with three LLMs and found that hypothesis generation is primarily driven by the model’s inherent priors.
Outcome: The proposed models generate high-quality hypotheses that can generalize to new instances when guided by in-context demonstrations.
EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis Generation (2026.acl-long)

Copied to clipboard

Challenge: Scientific discovery evolution does not occur ex nihilo but is characterized by structural deepening and reconfiguration of existing functionalities.
Approach: They propose a framework for hypothesis generation based on evolutionary narratives . they extract structured P-M-L-F quadruples from citation networks and introduce a mechanism to assess their semantic compatibility.
Outcome: The proposed framework reduces logical disconnects by evaluating its semantic compatibility.
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research.
Approach: They organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication.
Outcome: The authors summarize the current state of research in three main areas: hypothesis formulation, hypothesis validation, and manuscript publication.
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables (2025.findings-emnlp)

Copied to clipboard

Challenge: randomized experiments provide strong inferences, but are often infeasible due to ethical or practical constraints.
Approach: They propose a benchmark where the objective is to complete a partial causal graph.
Outcome: The proposed benchmarks show that they can hypothesize backdoor variables between a cause and its effect.
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System (2025.acl-long)

Copied to clipboard

Challenge: Recent AI methods have shown promise in tasks such as hypothesis generation and experimental design, but they fail to replicate the collaborative nature of real-world scientific practices.
Approach: They propose a virtual scientific system that mimics the collaborative nature of scientific research by organizing a team of agents to generate, evaluate, and refine research ideas.
Outcome: The proposed system outperforms the state-of-the-art method in producing new scientific ideas and offers valuable insights to guide future research.
OPINE: A Prior-calibrated Scoring Framework for LLM-based Multi-label Scientific Opinion Classification (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for scientific opinion classification rely on direct label generation and are limited by the multi-label nature of scientific expressions.
Approach: They propose a framework that reformulates scientific opinion classification as a controllable pipeline.
Outcome: The proposed framework outperforms baseline models on 18 discourse functions in micro, macro, and example settings.
Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for retrieval augmented generation are based on a simplistic binary choice of relying on external contexts or memory.
Approach: They propose a framework that transforms conflict resolution into an active deliberation process by incorporating contradictions as opportunities for deeper reasoning.
Outcome: Experiments show that DoT outperforms state-of-the-art methods while generating transparent debate transcripts that explain its decisions.
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Language models are accelerating scientific research by automating hypothesis generation and implementation.
Approach: They ask whether LMs can forecast the empirical success of research ideas before experiments . they frame evaluation as a reasoning task via Reinforcement Learning with Verifiable Rewards .
Outcome: The proposed model outperforms off-the-shelf models in 77.1% of the evaluations . the model outpersforms GPT-5 in the evaluation of 11,488 idea pairs .
GUIDE: Towards Scalable Advising for Research Ideas (2026.acl-long)

Copied to clipboard

Challenge: Existing systems that provide detailed, constructive feedback on academic papers struggle with review fidelity.
Approach: They explore factors that underlie the development of robust advising systems . large language models have shown remarkable progress in tasks from text generation to code synthesis .
Outcome: The proposed model outperforms general-purpose language models in acceptance rates for self-ranked top-30% submissions to ICLR 2025.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations